Read the emotion behind a two-minute video.
Record yourself answering five short prompts. Audio and facial motion are analysed together, then mapped onto four affective indicators — privately, in seconds.
- Modalities
- 2 Modalities
- Emotion classes
- 8 Emotion classes
- Training samples
- 2,452 Training samples
How it works
Three steps, one short recording
Nothing to install and nothing to fill in. Record, wait, read the breakdown.
Record
Answer five guided prompts on camera. The clip never leaves your account.
Analyze
Speech, facial motion and their fusion are scored by three trained networks.
Result
See the emotion distribution, indicator scores and how each modality voted.
Inside the model
Two networks, one fused verdict
A speech network and a 3D video network each predict an emotion on their own. Their features and probabilities are concatenated into a single 568-dimension vector that a fusion network turns into the final distribution.
Audio
CNN + BiLSTM + RNN over log-Mel spectrograms
Video
ResNet3D r3d_18 over 16-frame clips
Fusion
MLP over the 568-dim fused vector
From recording to indicator
-
1
Extract audio
ffmpeg → 16 kHz mono WAV
-
2
Score speech
128-band log-Mel → CNN + BiLSTM + RNN
-
3
Score facial motion
16 frames @ 112 px → ResNet3D r3d_18
-
4
Fuse both signals
568-dim vector → Fusion MLP
-
5
Map to indicators
Softmax → rule-based indicator scores
Run your first check
A two-minute recording is all the model needs.
Get StartedThis result is not a medical diagnosis. Please consult a professional.